Papers with text-to-speech synthesis

9 papers
Neural Text Normalization with Subword Units (N19-2)

Copied to clipboard

Challenge: Text normalization (TN) is an important step in conversational systems.
Approach: They frame text normalization as a machine translation task and tackle it with sequence-to-sequence models.
Outcome: The proposed model normalizes written text to its spoken form to facilitate speech recognition and text-to-speech synthesis.
Thesis Proposal: Development of End-to-End Speech Translation Models for Indian Languages (2026.eacl-srw)

Copied to clipboard

Challenge: Existing approaches to speech-to-speech translation rely on cascaded pipelines . current approaches rely only on text representations, but they suffer from errors and latency . a new direct speech translation framework is proposed to bridge linguistic gaps .
Approach: They propose a sequence-to-sequence direct speech translation framework that can translate speech from one Indian language to another without relying on intermediate text representations.
Outcome: The proposed framework can translate speech from one Indian language to another without relying on intermediate text representations.
Autoregressive Speech Synthesis without Vector Quantization (2025.acl-long)

Copied to clipboard

Challenge: MELLE is a novel language modeling approach for text-to-speech synthesis that generates continuous tokens from text . authors demonstrate that it reduces the need for vector quantization and improves model robustness .
Approach: They propose to autoregressively generate continuous mel-spectrogram frames directly from text condition, bypassing vector quantization.
Outcome: The proposed model achieves superior performance across multiple metrics and is more streamlined.
Improving homograph disambiguation with supervised machine learning (L18-1)

Copied to clipboard

Challenge: a new system for text-to-speech synthesis uses rule-based homograph disambiguation . a simple application of machine learning produces significant improvements in homograph ambiguity .
Approach: They propose a rule-based homograph disambiguation system for text-to-speech synthesis at Google . they compare it to a new system which performs disambiguations using classifiers trained on labeled data .
Outcome: The proposed system is more accurate than hand-written rules or machine learning alone.
A Challenge Set and Methods for Noun-Verb Ambiguity (D18-1)

Copied to clipboard

Challenge: English part-of-speech taggers make egregious errors related to noun-verb ambiguity, despite having achieved 97%+ accuracy on the WSJ Penn Treebank since 2002.
Approach: They propose to use a WSJ dataset to identify 30,000 examples of noun-verb ambiguity . they find that english part-of-speech taggers make egregious errors related to nouns and verbs .
Outcome: The proposed model improves on the WSJ Penn Treebank by 14% and 52% relative to the previous model.
Systematic Inequalities in Language Technology Performance across the World’s Languages (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages.
Approach: They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Outcome: The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Design and Development of Speech Corpora for Air Traffic Control Training (L18-1)

Copied to clipboard

Challenge: The current state-of-the-art training procedures involve retired pilots that train as virtual plane pilots and process the spoken prompts to form that can be entered into software that simulates the plane movement on the radar screen.
Approach: They describe the process of creating domain-specific speech corpora containing air traffic control (ATC) communication prompts.
Outcome: The proposed system could be used for training air traffic controllers in the Czech Republic.
Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks (2022.lrec-1)

Copied to clipboard

Challenge: Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining.
Approach: They propose to modify the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion.
Outcome: The proposed method improves the quality of the voice conversion system and the speaker similarity.
Revisiting Three Text-to-Speech Synthesis Experiments with a Web-Based Audience Response System (2024.lrec-main)

Copied to clipboard

Challenge: Audience Response System (ARS) evaluations are not well understood for text-to-speech synthesis (TTS) evaluation is a key weakness in the field and needs to adapt to be better-suited for this new generation of voices.
Approach: They revisit three published TTS studies and perform an ARS-based evaluation on the stimuli used in each study.
Outcome: The results show that Audience Response System (ARS) is highly useful for evaluating long and continuous stimuli.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations